New Multimodal fusion model — audio and video read together

Read the emotion behind a two-minute video.

Record yourself answering five short prompts. Audio and facial motion are analysed together, then mapped onto four affective indicators — privately, in seconds.

Modalities
2 Modalities
Emotion classes
8 Emotion classes
Training samples
2,452 Training samples

How it works

Three steps, one short recording

Nothing to install and nothing to fill in. Record, wait, read the breakdown.

01

Record

Answer five guided prompts on camera. The clip never leaves your account.

02

Analyze

Speech, facial motion and their fusion are scored by three trained networks.

03

Result

See the emotion distribution, indicator scores and how each modality voted.

Inside the model

Two networks, one fused verdict

A speech network and a 3D video network each predict an emotion on their own. Their features and probabilities are concatenated into a single 568-dimension vector that a fusion network turns into the final distribution.

Audio

CNN + BiLSTM + RNN over log-Mel spectrograms

Video

ResNet3D r3d_18 over 16-frame clips

Fusion

MLP over the 568-dim fused vector

Read the model card

From recording to indicator

  1. 1

    Extract audio

    ffmpeg → 16 kHz mono WAV

  2. 2

    Score speech

    128-band log-Mel → CNN + BiLSTM + RNN

  3. 3

    Score facial motion

    16 frames @ 112 px → ResNet3D r3d_18

  4. 4

    Fuse both signals

    568-dim vector → Fusion MLP

  5. 5

    Map to indicators

    Softmax → rule-based indicator scores

Run your first check

A two-minute recording is all the model needs.

Get Started

This result is not a medical diagnosis. Please consult a professional.